Papers with predictive models
Automatically Cataloging Scholarly Articles using Library of Congress Subject Headings (2021.eacl-srw)
Copied to clipboard
| Challenge: | Currently, nearly 40 institutions have registered their repositories with RAMP . manual cataloging of articles using LCSH is a challenge due to the rapid growth of articles . |
| Approach: | They propose to automatically annotate articles with Library of Congress Subject Headings . they use web scraping to extract keywords for a collection of articles from RAMP . |
| Outcome: | The proposed approach predicts LCSH for scholarly articles using keywords extracted from RAMP . the proposed model is validated by a multi-label classification problem. |
How Predictable is Your State? Leveraging Lexical and Contextual Information for Predicting Legislative Floor Action at the State Level (C18-1)
Copied to clipboard
| Challenge: | a study of state legislative initiatives shows that state legislatures have significant power over certain areas. |
| Approach: | They propose to use lexical content of over 1 million bills to build predictive models . they also use contextual legislature and legislator derived features to compare models based on state specific baselines . |
| Outcome: | The proposed models improve on baselines in all 50 states and D.C. lexical content, contextual features and legislative processes are used to build the models. |
Predicting Difficulty and Discrimination of Natural Language Questions (2022.acl-short)
Copied to clipboard
| Challenge: | Item Response Theory (IRT) has been used to numerically characterize question difficulty and discrimination for human subjects in domains including cognitive psychology and education. |
| Approach: | They explore the relationship between difficulty and discrimination in question-answering contexts by using IRT to characterize item difficulty and item discrimination. |
| Outcome: | The proposed models can predict difficulty and discrimination parameters for new questions and explain them with features of questions, answers, and associated contexts. |
Literature-Augmented Clinical Outcome Prediction (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing approaches to clinical outcome prediction use only clinical notes and general biomedical literature. |
| Approach: | They propose to retrieve patient-specific medical literature and incorporate it into predictive models by combining clinical notes with language models. |
| Outcome: | The proposed approach boosts predictive performance on three important clinical tasks in comparison to strong LM baselines, increasing F1 by up to 5 points and precision@Top-K by a large margin of over 25%. |
Predicting Foreign Language Usage from English-Only Social Media Posts (N18-2)
Copied to clipboard
| Challenge: | Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. |
| Approach: | They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets. |
| Outcome: | The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter. |
Bias Mitigation in Machine Translation Quality Estimation (2022.acl-long)
Copied to clipboard
| Challenge: | despite advances in machine translation, the accuracy and fluency of translations cannot be guaranteed without a reference translation. |
| Approach: | They propose to use auxiliary tasks to mitigate partial input bias . they aim to train a multitask architecture with an auxiliary binary classification task . |
| Outcome: | The proposed models reduce partial input bias while maintaining the overall performance. |
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)
Copied to clipboard
| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
Calibrating Zero-shot Cross-lingual (Un-)structured Predictions (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing need for model calibration when natural language models are deployed in critical tasks. |
| Approach: | They compare model calibration methods in a context of zero-shot cross-lingual transfer with pre-trained language models. |
| Outcome: | The proposed method fails to calibrate more complex confidence estimations in structured predictions compared to expressive alternatives like Gaussian Process Calibration. |
WEXEA: Wikipedia EXhaustive Entity Annotation (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for extracting factual knowledge from text are limited to a few subtasks. |
| Approach: | They propose to use Wikipedia to build a corpus with exhaustive annotations of entity mentions. |
| Outcome: | The proposed system can be used to build supervised datasets and can be reproduced by everyone. |
Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and Study (2024.lrec-main)
Copied to clipboard
| Challenge: | Mental illness can negatively impact individuals’ quality of life as it is considered one of the causes of years lived with disability and it is related to high suicide rates. |
| Approach: | They collect first dataset of textual posts by same users before and after being diagnosed with depression and build multiple predictive models based on Transformers and BERT. |
| Outcome: | The proposed model can be used to detect depression and suicidal thoughts in users who are not diagnosed with depression or suicide. |
Generating Realistic Natural Language Counterfactuals (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to explain ML tasks for natural language text are either unrealistic or introduce imperceptible changes. |
| Approach: | They propose a method that combines a conditional GAN and embeddings of a pretrained BERT encoder to model-agnostically generate realistic natural language text counterfactuals for explaining regression and classification tasks. |
| Outcome: | The proposed method outperforms baseline methods on fidelity and human judgments of naturalness across multiple datasets and multiple predictive models. |
Supporting Cognitive and Emotional Empathic Writing of Students (2021.acl-long)
Copied to clipboard
| Challenge: | Empathy skills are an elementary skill in society for daily interaction and professional communication and are therefore elementary for educational curricula. |
| Approach: | They propose an annotation approach to capture emotional and cognitive empathy in student-written peer reviews on business models in germany. |
| Outcome: | The proposed annotation scheme guides annotators to a substantial to moderate agreement with the model and shows that it is effective. |
Psycholinguistic Tripartite Graph Network for Personality Detection (2021.acl-long)
Copied to clipboard
| Challenge: | Existing work on personality detection from online posts adopts multifarious deep neural networks to represent the posts and builds predictive models in a data-driven manner without the exploitation of psycholinguistic knowledge. |
| Approach: | They propose a psycholinguistic knowledge-based tripartite graph network, TrigNet, which consists of a tripartitic graph network and a BERT-based graph initializer. |
| Outcome: | The proposed graph network outperforms the existing state-of-the-art model by 3.47 and 2.10 points in average F1 on two datasets. |
Modeling the Differential Prevalence of Online Supportive Interactions in Private Instant Messages of Adolescents (2025.findings-naacl)
Copied to clipboard
| Challenge: | Approximately two-thirds (68%) of American teenagers aged 13-17 have reported that social media make them feel as though they have people who will support them during challenging times. |
| Approach: | They propose to use the Social Support Behavioral Code to detect and model gender-based and pair-or-group disparities in online supportive interactions among adolescents. |
| Outcome: | The proposed model can be used to model gender-based and pair-or-group disparities in supportive interactions among adolescents. |
Multi Task Learning For Zero Shot Performance Prediction of Multilingual Models (2022.acl-long)
Copied to clipboard
| Challenge: | Massively Multilingual Transformer based Language Models have been shown to be effective on zero-shot transfer across languages, though performance varies from language to language depending on pivot language(s) used for fine-tuning. |
| Approach: | They propose to combine multi-task learning problems with multi-lingual Transformers to model zero-shot transfer across languages. |
| Outcome: | The proposed model can predict zero-shot transfer across languages with a multi-task learning problem with pretraining data in very few languages. |
Crowdsourcing and Validating Event-focused Emotion Corpora for German and English (P19-1)
Copied to clipboard
| Challenge: | Existing studies on automatic recognition of emotions in text have achieved promising results, but there is a shortage of resources for non-English languages, with few exceptions, like Chinese. |
| Approach: | They propose to use a crowdsourced German emotion corpus to build a corpus similar to the English ISEAR emotion dataset. |
| Outcome: | The proposed model performs well in German and English, but lacks the resources for non-English languages. |
LIFTED: Multimodal Clinical Trial Outcome Prediction via Large Language Models and Mixture-of-Experts (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Clinical trials are costly and pivotal processes that require substantial expenses . a new approach to integrate multimodal data for clinical outcome prediction is needed . |
| Approach: | a proposed framework transforms modality-specific data into natural language descriptions . a sparse Mixture-of-Experts mechanism then identifies shared patterns across modalities . |
| Outcome: | a proposed framework outperforms baseline methods in predicting clinical trial outcomes . it transforms modality-specific data into natural language descriptions, encoded via unified encoders . |
Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control (D19-1)
Copied to clipboard
| Challenge: | Selective rationalization is a common mechanism to ensure that predictive models reveal how they use any available features. |
| Approach: | They propose a co-operative method which uses introspection to explicitly predict and incorporate the outcome into the selection process. |
| Outcome: | The proposed model maintains high predictive accuracy and leads to comprehensive rationales. |
NLP for preserving Torlak, a vulnerable low-resource Slavic language (2025.coling-main)
Copied to clipboard
| Challenge: | Torlak is an endangered, low-resource Slavic language with a high degree of areal and inter-speaker variation. |
| Approach: | They aim to improve the prediction of morphosyntactic annotations for this low-resource Slavic language using the fine-tuning of large language models. |
| Outcome: | The proposed models improve the prediction of morphosyntactic annotations for Torlak using fine-tuning of large language models. |
Modeling Empathy and Distress in Reaction to News Stories (D18-1)
Copied to clipboard
| Challenge: | a recent work on empathy prediction has underestimated the complexity of the phenomenon and lacks a shared corpus. authors present a novel annotation methodology which reliably captures empathy assessments by the writer of a statement using multi-item scales. |
| Approach: | They propose a method which captures empathy assessments by the writer of a statement using multi-item scales. |
| Outcome: | The proposed method distinguishes between multiple forms of empathy, empathic concern, and personal distress, as recognized throughout psychology. |
The Role of Pragmatic and Discourse Context in Determining Argument Impact (D19-1)
Copied to clipboard
| Challenge: | Recent work shows that attributes of both the audience and communicator constitute important cues for determining argument strength. |
| Approach: | They propose to use a dataset to study the pragmatic and discourse context of argumentative claims to build predictive models that incorporate the pragmatic context of the argument. |
| Outcome: | The proposed models outperform models that rely on claim-specific linguistic features for predicting the perceived impact of individual claims within a particular line of argument. |
Evidence-guided Inference for Neutralized Zero-shot Transfer (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing knowledge transfer frameworks that use label skewness to neutralize biased language are costly and impractical when it comes to scarcely labeled data. |
| Approach: | They propose a neutralized Knowledge Transfer framework to equip pre-trained language models with neutralized transferability. |
| Outcome: | The proposed framework shows that it can be used to train pre-trained models with neutralized transferability . it is compared with baselines with a zero-shot cross-domain transfer setting . |
Modeling Persuasive Discourse to Adaptively Support Students’ Argumentative Writing (2022.acl-long)
Copied to clipboard
| Challenge: | Argumentation is an omnipresent rudiment of daily communication and thinking . humans struggle to develop argumentation skills due to a lack of individual and instant feedback in their learning process. |
| Approach: | They propose an argumentation annotation approach to model argumentative discourse in student-written business model pitches and embed it into an adaptive writing support system for students that provides individual argumentation feedback. |
| Outcome: | The proposed method annotates a corpus of 200 business model pitches in german and measures their self-efficacy and ease-of-use in a real-world writing exercise. |
Small Town or Metropolis? Analyzing the Relationship between Population Size and Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Prior studies have examined how location affects the type of language that people use . recent electoral results in the united states exemplify a divide in the political opinions of those living in densely populated areas . |
| Approach: | They analyze tweets from different Twitter users to determine whether they are from an urban or rural area. |
| Outcome: | The proposed model trains predictive models to predict whether a user is from an urban or rural area. |
Predictive Chemistry Augmented with Text Retrieval (2023.emnlp-main)
Copied to clipboard
| Challenge: | TextReact is a new method to augment predictive chemistry with text descriptions retrieved from the literature. |
| Approach: | They propose a method that directly augments predictive chemistry with texts retrieved from the literature. |
| Outcome: | The proposed method outperforms existing models trained on molecular data. |
Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party Debates (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on persuasion in online forums focuses on identifying debate winners and winning negotiation games. |
| Approach: | They adopt a hierarchical generative Variational Autoencoder model to model winning arguments . they propose competing hypotheses about the nature of argumentation . |
| Outcome: | The proposed model predicts winning arguments in reddit debates . it uses a hierarchical generative Variational Autoencoder to model argumentation . |
RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising results for mining EHRs . translating time-stamped sequences into plain text can obscure both temporal structure and code identities, weakening the ability to capture code co-occurrence and longitudinal regularities. |
| Approach: | They propose a time-aware LLM framework that integrates structured EHR encoders through prompt tuning without modifying underlying architectures. |
| Outcome: | Experiments on MIMIC-III and MIMIC IV show that RePrompT outperforms both EHR-based and LLM-based baselines across multiple clinical prediction tasks. |